In this lesson
Transformers and Attention
The architecture behind ChatGPT, BERT, and every modern language model. Understand self-attention, and you understand the engine driving the AI revolution.
Before Transformers: The Sequence Problem
Text is fundamentally different from an image. An image has fixed spatial structure: pixels at position (x, y) are always there. A sentence is a sequence of variable length, and the relationship between words can span long distances. In "The trophy did not fit in the bag because it was too large," what does "it" refer to? The trophy, not the bag. Understanding this requires relating "it" to "trophy" across eight words of distance.
For years, the dominant approach to sequence data was Recurrent Neural Networks (RNNs). An RNN processes a sequence one element at a time, passing a "hidden state" from step to step. The hidden state at step 50 is supposed to remember what happened at step 1. The problem is it often does not. Information gets diluted as it passes through many sequential multiplications.
Long Short-Term Memory networks (LSTMs), introduced by Hochreiter and Schmidhuber in 1997, added gating mechanisms that helped retain important information over longer distances. They powered breakthrough results in machine translation and speech recognition through the 2010s. But LSTMs still had a fundamental bottleneck: the entire history of a sequence had to be compressed into a single fixed-size hidden state vector before generating each output token.
In 2015, Bahdanau, Cho, and Bengio introduced the first attention mechanism for sequence-to-sequence models, allowing the decoder to look directly at all encoder states rather than relying solely on the compressed hidden state. This was enormously effective for machine translation. It was also the seed of a more radical idea: what if the entire architecture was built around attention, with no recurrence at all?
RNNs process sequences step by step: step 2 cannot begin until step 1 is complete. This sequential dependency means you cannot parallelise the computation across the length of the sequence. On modern GPUs and TPUs, which are designed for massively parallel computation, this is a serious limitation. Transformers replace sequential recurrence with parallel attention, processing all positions simultaneously, which is a large part of why they scaled so dramatically.
The Core Idea: Attention
The intuition behind attention is something humans do naturally when reading. When you read the sentence "The scientist published the paper after she reviewed it carefully," and you want to understand what "she" refers to, your brain does not re-read the entire sentence. It pulls out the relevant word "scientist" and connects them. You are paying attention to specific parts of the input based on relevance to the current task.
In a Transformer, every token in a sequence is given the ability to directly attend to every other token, computing a weighted combination that reflects how relevant each position is. If "she" strongly attends to "scientist," the representation of "she" will be heavily influenced by the representation of "scientist," correctly capturing the coreference relationship.
The mechanism uses a query-key-value framework drawn from information retrieval. Think of it like a soft database lookup:
Query (Q)
What the current token is "looking for." Each token generates a query vector that represents what kind of information it needs from the context.
Key (K)
What each token "offers." Every token generates a key vector that describes the kind of information it contains. A high Query-Key match means high relevance.
Value (V)
The actual content contributed. If a token's key matches the query well, its value vector contributes heavily to the output. The output is a weighted average of all values.
Q, K, and V are not separate inputs. They are all derived from the same token embeddings through three different learned linear projections. This means the model learns what to look for, what to advertise, and what to contribute, all from the same underlying representations.
Self-Attention Step by Step
Let us walk through the exact computation. Given a sequence of tokens, each represented as a vector of dimension d_model:
Step 1: Project into Q, K, V
Multiply each token embedding by three learned weight matrices (W_Q, W_K, W_V) to get three vectors per token: a query Q, a key K, and a value V. The dimensions of Q and K must match (call it d_k) since we will compute their dot product.
Step 2: Compute Attention Scores
For each token acting as a query, compute a dot product with every token's key. A high dot product means the query and key are aligned, that is, this pair of tokens has a strong relationship.
Step 3: Scale and Softmax
Divide each score by the square root of d_k. This scaling prevents the dot products from growing too large in magnitude (which would push the softmax into regions where gradients become tiny). Then apply softmax to get a probability distribution over all positions, the attention weights.
Step 4: Weighted Sum of Values
Multiply the attention weights by the value vectors and sum. The output for each token is a weighted mixture of all value vectors, where the weights reflect how relevant each position was. Tokens that scored highly in the attention contribute more to the output.
An illustrative self-attention weight matrix for a 4-token sentence. "gone" attends most strongly to itself (0.45) and to "cat" (0.35), capturing the semantic relationship. Every row is a softmax distribution over all key positions.
A Concrete Implementation
The mathematics of scaled dot-product attention can be implemented in about ten lines of Python:
import numpy as np def softmax(x): # Subtract max for numerical stability — prevents overflow in exp() exp_x = np.exp(x - np.max(x, axis=-1, keepdims=True)) return exp_x / exp_x.sum(axis=-1, keepdims=True) def scaled_dot_product_attention(Q, K, V): """ Q: (seq_len, d_k) — queries K: (seq_len, d_k) — keys V: (seq_len, d_v) — values """ d_k = Q.shape[-1] # Step 1: compute all pairwise scores — (seq_len, seq_len) scores = Q @ K.T / np.sqrt(d_k) # Step 2: apply softmax row-wise to get attention weights weights = softmax(scores) # Step 3: weighted sum of value vectors output = weights @ V return output, weights # Example: 4 tokens, d_k = 4 np.random.seed(42) seq_len, d_k = 4, 4 Q = np.random.randn(seq_len, d_k) K = np.random.randn(seq_len, d_k) V = np.random.randn(seq_len, d_k) output, attn_weights = scaled_dot_product_attention(Q, K, V) print("Attention weights (each row sums to 1):") print(attn_weights.round(3)) print(f"\nOutput shape: {output.shape}")
This is the complete core of self-attention. Each row in the weight matrix is a distribution over all four positions. The output for each token is a weighted blend of all value vectors according to these weights.
Multi-Head Attention
A single attention head learns one way to relate tokens. But a sentence has multiple types of relationships simultaneously. "John gave Mary the book" contains relationships between giver and receiver, between action and object, between subject and verb. A single attention pattern cannot capture all of these at once.
Multi-head attention runs several attention operations in parallel, each with its own independent W_Q, W_K, and W_V projection matrices. Each "head" is free to learn a different type of relationship. The outputs from all heads are concatenated and projected through a final linear layer.
The original Transformer used 8 heads with d_model = 512. Each head operates in a reduced dimension of 512 / 8 = 64, so the total computational cost is comparable to a single head at full dimension. Research on analysing what individual heads learn has found that some heads specialise in syntactic relationships (subject-verb agreement), some in coreference (pronoun-noun), and some in positional patterns, though this specialisation is not guaranteed or required by the design.
MultiHead(Q, K, V) = Concat(head_1, ..., head_h) W_O, where each head_i = Attention(Q W_Q_i, K W_K_i, V W_V_i). The W_O matrix projects the concatenated output back to d_model dimensions.
The Transformer Block
Self-attention alone is not a complete model. The full Transformer block wraps it with several additional components, each serving a specific engineering purpose.
One Transformer encoder block. The green dashed lines are residual (skip) connections, exactly as in ResNet. They help gradients flow through deep stacks. Layer Norm stabilises training. The feed-forward network processes each position independently after attention has aggregated context.
Residual Connections
The output of each sub-layer is added to its input before normalisation. This is the same skip connection idea from ResNet (He et al., 2015), and it is critical for training deep stacks of Transformer blocks without vanishing gradients.
Layer Normalisation
Introduced by Ba, Kiros, and Hinton in 2016, Layer Norm normalises across the feature dimension for each token independently. Unlike Batch Norm, it does not depend on batch size, making it better suited to variable-length sequences.
Feed-Forward Network
After attention aggregates context, a two-layer MLP (Dense → ReLU → Dense) processes each token independently. This is where much of the model's "knowledge" is stored. In the original paper, the inner dimension d_ff = 2,048 (four times d_model = 512).
Positional Encoding
Attention has no inherent notion of order. To tell the model where each token sits in the sequence, positional information is added to the embeddings. The original paper used fixed sinusoidal functions. Modern models (GPT, BERT variants) typically use learned positional embeddings.
The 2017 Breakthrough: "Attention Is All You Need"
The Transformer architecture was introduced in the paper "Attention Is All You Need" by Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, all researchers at Google at the time. It was presented at NeurIPS 2017.
The paper's central claim was bold: recurrence and convolution were not necessary for sequence modelling. Attention alone, applied in parallel across all positions, was sufficient and indeed superior. The model they proposed achieved state-of-the-art results on English-to-German and English-to-French machine translation tasks, while training far faster than the recurrent models of the time.
The base Transformer had these specs:
# Original Transformer "base" hyperparameters (Vaswani et al., 2017) d_model = 512 # embedding / model dimension d_ff = 2048 # feed-forward inner dimension (4× d_model) num_heads = 8 # attention heads per block d_k = d_v = 64 # dimension per head (512 / 8) num_layers = 6 # number of encoder blocks (and 6 decoder blocks) dropout = 0.1 # dropout rate applied throughout # Total parameters: approximately 65 million
The paper also introduced the concept of scaled dot-product attention. Removing the scaling factor 1/sqrt(d_k) causes the dot products to grow large in high dimensions, pushing softmax into saturation regions where gradients nearly vanish. The scaling was a practical and important engineering decision.
The title is a direct challenge to the field's assumptions. Attention mechanisms had been used as add-ons to RNNs. The claim that attention alone, without any recurrence, was not just sufficient but better was a paradigm shift. The paper has been cited well over 100,000 times and spawned an entire new era of AI architecture design.
BERT and GPT: Two Strategies, One Architecture
The Transformer has an encoder stack and a decoder stack. Different research groups discovered that pre-training just one of these components on massive text datasets, then fine-tuning on specific tasks, produced remarkably capable models. Two paradigms emerged.
BERT (Encoder-Only)
Bidirectional Encoder Representations from Transformers. Google, 2018 (Devlin, Chang, Lee, Toutanova).
- Uses only the encoder stack
- Each token attends to all tokens in both directions simultaneously (bidirectional)
- Pre-trained with Masked Language Modelling: randomly mask 15% of tokens, predict them from context
- Best for: understanding tasks — question answering, sentiment analysis, named entity recognition, text classification
- Not designed to generate text
GPT (Decoder-Only)
Generative Pre-trained Transformer. OpenAI, 2018 (Radford, Narasimhan, Salimans, Sutskever).
- Uses only the decoder self-attention with causal masking
- Each token can only attend to previous tokens (left-to-right only)
- Pre-trained with Causal Language Modelling: predict the next token given all previous tokens
- Best for: generation tasks — writing, summarisation, conversation, code
- ChatGPT, GPT-4, and most modern chatbots use this paradigm
The Scale-Up Story
In a decoder-only model, when predicting token 5, the model cannot look at tokens 6, 7, 8... (they do not exist yet in generation). This is enforced by a causal mask: a triangular mask that sets attention scores to negative infinity for all future positions before softmax, making their weights effectively zero. During training, the model processes all positions in parallel but each position only "sees" its left context, mimicking the sequential generation process efficiently.
Using Transformers Today: Hugging Face
Training a Transformer from scratch requires enormous data and compute. In practice, you almost never do this. Instead, you use pre-trained models through the Hugging Face transformers library, the dominant open-source platform for working with Transformer models. It provides a unified API for hundreds of pre-trained models in dozens of languages.
Run this once in your terminal or a Colab cell: pip install transformers torch. The library handles downloading model weights automatically on first use.
from transformers import pipeline # ── 1. Sentiment Analysis using a fine-tuned BERT model ─────────────── # Default model: distilbert-base-uncased-finetuned-sst-2-english # (a compressed BERT fine-tuned on Stanford Sentiment Treebank) classifier = pipeline("sentiment-analysis") texts = [ "I really enjoyed this course, the explanations are excellent!", "This was confusing and poorly structured.", "Convolutional neural networks are used in image recognition." ] results = classifier(texts) for text, result in zip(texts, results): print(f"Text: '{text[:50]}...'") print(f" Label: {result['label']}, Score: {result['score']:.4f}\n") # ── 2. Text Generation using GPT-2 ──────────────────────────────────── generator = pipeline("text-generation", model="gpt2") prompt = "Artificial intelligence is transforming" output = generator( prompt, max_new_tokens=50, # generate 50 new tokens after the prompt num_return_sequences=1, do_sample=True, # sample from the distribution (not greedy) temperature=0.8 # lower = more predictable, higher = more creative ) print("Generated text:") print(output[0]['generated_text']) # ── 3. Question Answering using BERT ────────────────────────────────── qa = pipeline("question-answering") context = """ The Transformer architecture was introduced in 2017 by Vaswani et al. in the paper 'Attention Is All You Need'. It relies entirely on self-attention mechanisms and has become the foundation of modern language models including BERT and GPT. """ result = qa(question="When was the Transformer introduced?", context=context) print(f"\nQuestion: When was the Transformer introduced?") print(f"Answer: {result['answer']} (confidence: {result['score']:.4f})")
What Just Happened?
The pipeline function downloaded a pre-trained model (several hundred MB), loaded its weights, and ran inference. For sentiment analysis, it used a DistilBERT model fine-tuned on human-labelled movie and review data. For question answering, it used a BERT model trained to extract spans from context paragraphs. None of the heavy lifting required any data or training from us.
These models were pre-trained on billions of words of text, which gave them broad language understanding. They were then fine-tuned on labelled examples specific to each task. This pre-train-then-fine-tune paradigm is covered in depth in Lesson 4.5 on Transfer Learning. It is the reason a beginner can achieve expert-level NLP results with three lines of code.
Key Takeaways
- Recurrent networks process sequences step by step and struggle to maintain context over long distances. Transformers process all positions in parallel using attention.
- Self-attention lets every token attend to every other token in a sequence. The attention weight between two tokens reflects how relevant one is to the other.
- Scaled dot-product attention computes Q and K dot products, scales by 1/sqrt(d_k) to prevent large magnitudes, applies softmax to get weights, then takes a weighted sum of V vectors.
- Multi-head attention runs several attention heads in parallel. Each head can learn a different type of relationship (syntactic, semantic, positional). Outputs are concatenated and projected.
- A Transformer block wraps attention with a feed-forward network, residual connections, and Layer Normalisation. Blocks are stacked N times (6 in the original paper).
- Positional encoding adds sequence position information to token embeddings, since attention alone has no sense of word order.
- "Attention Is All You Need" (Vaswani et al., 2017, NeurIPS) introduced the architecture and showed recurrence was not needed for top-quality translation.
- BERT (encoder-only, bidirectional) excels at understanding tasks. GPT (decoder-only, causal) is designed for text generation. ChatGPT, GPT-4, and Claude are all GPT-style decoder models.
- The Hugging Face library lets you use pre-trained Transformers for dozens of NLP tasks with just a few lines of Python.
In Lesson 4.5, you will see how to take a pre-trained model and adapt it to your own specific task with a small amount of data. This technique, called fine-tuning, is how most real-world AI applications are built today. You do not need to train a model from scratch. You start from a powerful foundation and teach it your specific domain.